Papers with multilingual vision-language models
Stop Pre-Training: Adapt Visual-Language Models to Unseen Languages (2023.acl-short)
Copied to clipboard
| Challenge: | Existing studies have shown that the pre-training in English does not transfer well to other languages in a zero-shot setting. |
| Approach: | They propose a simple yet efficient approach to adapt VLP to unseen languages using MPLM. |
| Outcome: | The proposed approach outperforms state-of-the-art models without large parallel corpora across three tasks. |
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing models that account for perceptual differences in image captions are limited to use in English . culture-based tasks such as recognition, detection, and image retrieval are hindered by relying on English supervision. |
| Approach: | They propose and evaluate caption augmentation strategies to address these gaps . they use captions from german perception and captions that have been machine-translated or human-transcribed from English into german . |
| Outcome: | The proposed models achieve a mean recall improvement of +1.3, but still lack flexibility . cultural differences present in language with respect to object specificity and importance . |